Papers with human reliability
Learning and Evaluating Emotion Lexicons for 91 Languages (2020.acl-main)
Copied to clipboard
| Challenge: | Emotion lexicons describe the affective meaning of words but are limited in coverage for most languages. |
| Approach: | They propose a method for creating arbitrarily large emotion lexicons for any target language. |
| Outcome: | The proposed method exceeds human reliability for some languages and variables. |
Why Do LLM-based Web Agents Fail? A Hierarchical Planning Perspective (2026.acl-long)
Copied to clipboard
| Challenge: | Existing evaluations focus primarily on end-to-end success, offering limited insight into where failures arise. |
| Approach: | They propose a hierarchical planning framework that analyzes web agents across three layers . they show that structured Planning Domain Definition Language (PDDL) plans produce more concise and goal-directed strategies than natural language (NL) plans . |
| Outcome: | The proposed framework analyzes web agents across three layers to improve reasoning, grounding, and recovery. |